Back

Journal of Molecular Biology

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Journal of Molecular Biology's content profile, based on 232 papers previously published here. The average preprint has a 0.14% match score for this journal, so anything above that is already an above-average fit.

1
Quantifying the binding affinity of a pharmacological chaperone to transient unfolded states of a normally folded protein

Patra, S.; Garen, C. R.; Woodside, M. T.

2026-08-28 biophysics 10.64898/2026.08.27.747683 medRxiv
Top 0.2%
8.0%
Show abstract

Binding of ligands to partially or fully unfolded proteins can play a key role in the mechanism of cellular and pharmacological chaperones, facilitating proper folding. However, it is challenging to quantify the binding affinity of ligands for unfolded states in a protein that is normally folded, as the methods standardly used to destabilize the native fold also affect ligand binding. We used single-molecule force spectroscopy to unfold single protein molecules without altering solution condi-tions and observe interactions of a ligand with unfolded states. Focusing on pentosan polysulfate (PPS), an anti-prion pharmacological chaperone previously shown to interact with both the native and partially or fully unfolded states of the prion protein, we measured the concentration-dependent effects of PPS binding on the conformational dynamics of bank vole prion protein (BvPrP) molecules held in optical tweezers. We found that PPS stabilized certain partially unfolded intermediate states of BvPrP as well as the fully unfolded state. Strikingly, the tendency to bind unfolded states instead of the folded state increased as the PPS concentration was reduced, implying a higher affinity to unfolded states. From the relative amount of binding to unfolded versus folded states, we estimated that PPS bound roughly 100-fold more tightly to unfolded states than to the native state of PrP. These results reinforce the likely importance of unfolded states in prion misfolding and propagation. More generally, they show how binding affinity to transient, unstable states can be estimated.

2
RheoScale 2.0: Revealing the Hidden Roles of Protein Positions via Substitution Patterns

Liu, D.; Sreenivasan, S.; Gray, C. J.; Cleveland, H. C.; Swint-Kruse, L.

2026-08-11 biochemistry 10.64898/2026.08.10.743964 medRxiv
Top 0.2%
8.0%
Show abstract

A central challenge in molecular biology is understanding how amino acid substitutions modulate various features of protein function and stability. To illuminate the complexities of this relationship, high-throughput (HTP) assays are increasingly used to assess site-saturating mutagenesis libraries. A common downstream analysis is to average the set of twenty outcomes at each amino acid position for comparison with structural and evolutionary features. Average values clearly identify positions that tolerate most substitutions (neutral positions) and positions where most substitutions abolish activity (toggle positions). However, average values conceal the existence of rheostat positions, where different amino acid substitutions sample a wide range of outcomes. To quantitatively identify rheostat positions, we previously developed a histogram-based analysis that we here expand by: (i) incorporating new position classes observed in experimental studies of rheostat positions; (ii) formalizing a hierarchy of class assignments; (iii) refining error-based identification of neutral positions; and (iv) statistically assessing the robustness of class assignments to changes in experimental and computational parameters. RheoScale 2.0 is implemented in Excel and newly implemented in Python for facile integration with existing HTP pipelines; all parameters are customizable. Example analyses are shown for three HTP datasets of the SARS-CoV-2 papain-like protease. Results illustrate two aspects that influence interpretation of HTP data: First, position assignments (and substitution outcomes) depend highly on the measured feature. Second, many protein positions play multiple roles in the sequence-structure-function relationship. The recognition of varied position roles will advance understanding of pathogen evolution, protein engineering, and variant interpretation for personalized medicine. SummaryRheoScale 2.0 improves how high-throughput mutational data are interpreted by identifying protein positions where amino acid substitutions act like biological dimmer switches. By enabling more nuanced assignment of position behavior, beyond neutral or deleterious outcomes, this analysis framework advances studies of sequence-structure-function relationships and has broad relevance for understanding protein evolution, engineering proteins with desired properties, and interpreting variants linked to human disease. SOFTWARE AVAILABILITYhttps://github.com/liskinsk/RheoScale-calculator

3
Scop3P-Toolkit: executable structure-aware workflows linking PTMs, peptides, and mutations to protein function

Diaz, A.; Tichshenko, N.; Depoortere, B. G. J.; Andrade Buono, R.; De Geest, P.; Vranken, W. F.; Martens, L.; Ramasamy, P.

2026-08-09 bioinformatics 10.64898/2026.08.04.742789 medRxiv
Top 0.2%
8.0%
Show abstract

Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein-protein, protein-ligand, and host-pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voila applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.

4
ResiRuler: A Toolkit for Visualizing Residue-Residue Distances and Structural Changes in Biomolecular Models

Baker, T. H.; Ohi, M. D.; Salmen, W.

2026-08-23 bioinformatics 10.64898/2026.08.19.745761 medRxiv
Top 0.2%
7.2%
Show abstract

Proteins and their associated complexes often adopt multiple conformations, with the transitions between these states playing a critical role in biological function. However, the resulting structural heterogeneity can be challenging to visualize and communicate, often requiring manual inspection and time-consuming annotation of biomolecular structures. To address this, we developed ResiRuler, a local, browser-based tool that uses inter-residue distance measurements to quickly quantify atomic displacements and map changes in internal geometry across ensembles of related protein structures. By converting structural differences into residue-pair distance changes, ResiRuler enables rapid identification of regions undergoing coordinated motion, local rearrangement, or large-scale conformational change. The resulting visualizations can be exported as scripts for PyMOL and ChimeraX, allowing users to explore conformational differences and generate publication-quality molecular figures in their preferred visualization environment. Using atomic models in Macromolecular Crystallographic Information File (mmCIF) file format, ResiRuler aligns multiple structures and measures structural variation across models facilitating visualization and presentation of these differences. This allows for rapid visualization of which regions of proteins change among ensembles of structures. The program is available for download at https://github.com/tbaker67/ResiRuler on macOS and Linux operating systems.

5
Resolution-standardized evaluation of ligand atomic coordinates in crystallographic structures using machine learning

Miyaguchi, I.; Hata, H.; Kuribayashi, T.; Takahashi, S.; Kashima, A.; Murasaki, K.; Matsumoto, S.; Terayama, K.; Ohta, M.; Ikeguchi, M.

2026-08-20 molecular biology 10.64898/2026.08.17.745351 medRxiv
Top 0.2%
7.2%
Show abstract

Accurate assessment of ligand coordinate-density consistency across different resolutions remains challenging in macromolecular crystallography. We introduce the atomic Box Correlation Coefficient (aBCC), an atom-level metric for evaluating the consistency between ligand atomic coordinates and electron density in a resolution-standardized framework. To predict aBCC values from electron-density maps, we developed QAEmap, a machine-learning model based on three-dimensional convolutional neural networks (3D-CNNs). The model was trained using Fourier-truncated electron-density maps and corresponding ligand coordinates generated from high-resolution structures in the Protein Data Bank. It was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures. was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures.The prediction accuracy gradually decreased with decreasing resolution, but remained reliable up to [~]3.5 [A]. These results demonstrate that aBCC enables resolution-standardized atom-wise evaluation of coordinate-density consistency across different resolutions and provide a foundation for further development and refinement of machine learning-based coordinate validation. SynopsisWe introduce the atomic box correlation coefficient (aBCC), a machine learning-based metric for the resolution-standardized atom-level evaluation of ligand coordinate-density consistency in crystallographic structures. aBCC provides a common framework for assessing and communicating the local coordinate reliability between structural biologists and researchers in structure-based drug discovery.

6
Two heads are better than one: Single stranded DNA translocation of UvrD-family dimers vs. monomers

Mersch, K. N.; Nguyen, B.; Kozlov, A. G.; Lohman, T. M.

2026-08-07 biophysics 10.64898/2026.08.07.743547 medRxiv
Top 0.2%
7.2%
Show abstract

UvrD-family Superfamily 1A helicases are processive ATP-dependent motor proteins that function during DNA replication, recombination, repair, and transcription. UvrD-family monomers translocate along single stranded (ss) DNA with 3-to-5 directionality but must be activated by dimerization to become helicases in the absence of force or accessory factors. Mycobacterium tuberculosis (Mtb) UvrD1 helicase forms dimers via a disulfide bond between native cysteines in the 2B sub-domains of each monomer. E. coli UvrD forms non-covalent dimers using the same 2B domain interface as in Mtb UvrD1. Using both ensemble and single DNA molecule approaches we examined an E. coli UvrD variant (R421C), which forms covalent dimers with constitutive helicase activity. For the first time this has enabled us to compare the ssDNA translocation and helicase activities of covalent dimers and monomers. Crosslinked UvrD dimers exhibit much higher ssDNA translocation processivities than monomers, although with similar translocation rates. Crosslinked UvrD dimers also show highly processive DNA unwinding of thousands of base pairs, much higher than non-crosslinked UvrD dimers, while monomers show no DNA unwinding activity. DNA unwinding rates of crosslinked UvrD dimers are only [~]20% slower than ssDNA translocation rates, indicating they are "active" helicases that directly facilitate duplex destabilization.

7
Safety First: Input Screening for Protein Design Tools

Palmer, P.; Teran, N.; Wheeler, N.; Yassif, J. M.

2026-08-07 synthetic biology 10.64898/2026.08.04.740855 medRxiv
Top 0.2%
7.1%
Show abstract

As biological AI models become more powerful, practical biosecurity approaches are needed to support beneficial applications while reducing misuse risks. Sequence-similarity-based screening approaches are no longer adequate to safeguard biological AI models because these models can design molecules with novel sequences and structures. Therefore, a screening approach that takes function into account is needed. To address this need, we propose a new screening method for AI-enabled protein binder design tools. Our framework screens protein binding targets, with a focus on the human proteome, as opposed to the binder molecule itself. We constructed a database of 14,541 potentially harmful proteoform targets from the human proteome (7.1% of all human protein proteoforms) classified by biosecurity risk level. To discern structural and functional features, we evaluated constructs with an embedding-based screening method using the ESM-C protein language model. ESM-C achieved high accuracy for detecting variants of known targets (F1 scores >97%), with performance similar to BLASTP. However, ESM-C proved to be more effective at capturing functional relationships, distinguishing benign mutations from damaging ones where BLASTP did not. To characterize how screening would affect bioscience research, we measured flagging rates across diverse protein datasets. Flagging rates were significant for mammalian proteins weighted by publication frequency (23% for human, 20% for mouse), and rates for organisms distantly related to humans were minimal (<1.1% for bacteria, fungi, plants, and viruses). Among commercially relevant targets, 63% of antibody patent targets were classified as dual-use, reflecting that therapeutically important proteins often perform critical biological functions. To identify and flag risky user requests from protein binder design tools without placing an undue burden on scientific research and innovation, it will be essential to deploy this screening approach in a way that addresses the overlap our analysis showed between targets of concern and therapeutic targets-possibly in concert with tiered trusted access frameworks. This new method provides a foundation for proportionate safeguards for biological AI models that reduce misuse risks while preserving their benefits for legitimate research and demonstrates a concrete proof of principle that can be generalized to other protein design tools and biological AI models.

8
Computational Structural Analysis of POLG Variants R627Q and W748S with Model-Variability Controls

Friedl, A.; Manst, D.

2026-08-27 biophysics 10.64898/2026.08.25.747108 medRxiv
Top 0.2%
6.7%
Show abstract

Background: Comparisons between independently predicted wild-type and missense-variant protein structures can generate mechanistic hypotheses, but small apparent differences may reflect model-selection variability rather than mutation-specific effects. Methods: Human mitochondrial DNA polymerase gamma (POLG; UniProt P54098) variants p.Arg627Gln (R627Q) and p.Trp748Ser (W748S) were evaluated using five AlphaFold2-PTM network-model outputs per condition generated with one random seed under matched ColabFold settings. Ten pairwise wild type comparisons at each site described between-network model-selection variability. Variant effects were summarized across five within-network wild-type-versus-variant comparisons using rotation-invariant local C-alpha pair distances and local displacement after global and local alignment. Because these comparison designs differ, the wild-type distribution was used as context rather than a mutation-effect null. Wild-type cryo-EM structure 9GGF was used for contact and interface mapping. Experimental A467T and G848S structures 9GGE and 9GGC provided contextual benchmarks. Results: R627Q measurements fell within the range of between-network wild-type differences: its median mean local pair-distance change was 0.170 angstrom, compared with a wild-type median of 0.170 angstrom, and its locally aligned displacement was 0.265 versus 0.248 angstrom. W748S showed higher median values (0.168 versus 0.132 angstrom for pair-distance change; 0.236 versus 0.182 angstrom for locally aligned displacement), but the ranges overlapped and the comparison-design asymmetry precluded a calibrated mutation-effect percentile. Experimental A467T and G848S comparisons produced local changes of similar magnitude. In 9GGF, R627 and W748 directly shared a local microenvironment, with a minimum heavy-atom distance of 3.53 angstrom. R627 also formed short polar-contact candidates with D629 and D743, whereas W748 occupied a hydrophobic packing environment containing Y622 and F750. Both sites were more than 18 angstrom from nucleic acid, more than 30 angstrom from POLG2, and more than 33 angstrom from PZL-A in a ligand-bound structure. Conclusions: Available AlphaFold2 comparisons do not establish a mutation-specific structural deformation for either variant. Experimental-structure mapping supports testable physicochemical hypotheses involving a shared R627-W748 microenvironment - loss of an arginine-centered polar network for R627Q and disruption of a buried aromatic environment for W748S - but not direct DNA, POLG2, or PZL-A contact mechanisms. Matched control substitutions and independent seeds are required to calibrate small mutation-associated structural deltas.

9
Foldseek-Interface reveals a protein interface universe far from complete

Strom, J. M.; Cha, S.; Kim, R. S.; Sajal, H.; Gilchrist, C. L.; Steinegger, M.; Luck, K.

2026-08-25 bioinformatics 10.64898/2026.08.24.746585 medRxiv
Top 0.3%
6.5%
Show abstract

Protein-protein interactions mediate a vast range of cellular functions, requiring diverse modes of binding. While recent years have seen major efforts to chart and classify the protein structure universe, we lack comparable methods to assess and cluster that diversity in interface structure at interactome scale. Here, we present Foldseek-Interface, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces. It matches the accuracy of state-of-the-art tools while running up to 230 times faster. Applying it to all biological assemblies in the PDB, we cluster 3.1 million dimers into 77{,}167 interface clusters and use this resource to characterise interface diversity, evolution, and pathogen mimicry. Application of Foldseek-Interface to resources of predicted protein complex structures rapidly revealed putatively novel interface types worth further experimental interrogation. Foldseek-Interface and the interface cluster resource are freely available as webservers for search https://search.foldseek.com/interface and exploration https://interface.foldseek.com.

10
The conserved β-hairpin of the SUI1 domain is a dual-function structural module governing translation initiation and ribosome recycling in yeast

Zamyatnina, K. A.; Urakov, V. N.; Volynkina, I. A.; Stolboushkina, E. A.; Gerasimov, E. S.; Kats, L. M.; Kushnirov, V. V.; Kamenski, P. A.; Dmitriev, S. E.

2026-08-21 molecular biology 10.64898/2026.08.14.744980 medRxiv
Top 0.3%
6.1%
Show abstract

Most eukaryotic mRNAs encode a single functional polypeptide. Following translation termination, both the large and small ribosomal subunits are typically released from the mRNA by ribosome recycling factors. However, after translating short upstream open reading frames (uORFs) within the 5 untranslated regions (UTRs), ribosomes can remain associated with the mRNA and reinitiate translation. This process is regulated by the heterodimer MCTS1*DENR (Tma20p*Tma22p in yeast). DENR/Tma22p harbors a SUI1 domain, structurally homologous to the translation initiation factor eIF1/Sui1p, which features a conserved, positively charged {beta}-hairpin loop critical for eIF1 function. Despite this structural similarity, the functional significance of specific elements within DENR/Tma22p remains unexplored. Here, we used in vivo reporter assays in Saccharomyces cerevisiae to quantify reinitiation efficiency following translation of either a short uORF (in the 5 UTR) or a full-length coding sequence (in the 3 UTR). Systematic analysis of single, double, and triple deletions of TMA20, TMA22, and TMA64 (a homolog of Tma20p*Tma22p) revealed that the Tma20p*Tma22p complex exerts a dominant role over Tma64p in modulating reinitiation, while exhibiting functional interplay between the two factors. Using knockout strains complemented with Tma22p variants, we further demonstrated that the positively charged residues of the {beta}-hairpin loop 1 are essential for Tma22p recycling activity. Unexpectedly, deletion of the entire SUI1 domain was less deleterious, and eIF1/Sui1p was able to partially substitute for the SUI1 domain of Tma22p within a chimeric protein context. Our findings establish the {beta}-hairpin loop 1 of the DENR/Tma22p SUI1 domain as a critical determinant for ribosome recycling and reinitiation, and raise the question of whether MCTS1/Tma20p can promiscuously operate with both DENR/Tma22p and eIF1/Sui1p - two specialized factors that evolved from a common structural scaffold to govern distinct steps in the translation cycle.

11
From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

qin, y.; Pang, J.; Zhang, X.

2026-09-01 bioinformatics 10.64898/2026.08.26.747436 medRxiv
Top 0.3%
6.1%
Show abstract

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

12
Structural and functional basis of the non-canonical human Dicer-tRNA complex

Di Fazio, A.; Hirschi, S.; Battistini, F.; Santos, N.; Boot, J.; Ajit, K.; Abdullah, A.; Alagia, A.; Orozco, M.; Gullerova, M.

2026-08-13 molecular biology 10.64898/2026.08.12.744379 medRxiv
Top 0.4%
5.5%
Show abstract

Human Dicer (hDicer) is a key enzyme in the RNA interference (RNAi) pathway that generates [~]21-22 nt micro-RNA (miRNAs) and small interfering RNAs (siRNAs). We have previously shown that hDicer also generates tRNA-derived small RNAs (tsRNAs), which mediate nuclear gene silencing and regulate hundreds of disease-associated genes. As powerful and evolutionarily conserved cellular regulators, tsRNAs emerged as an important class of small RNAs. Therefore, it is essential to understand their biogenesis. However, the molecular and structural basis of tRNA cleavage by hDicer, as well as the role of chemical modifications such as 5-methylcytosine (m5C), in this process, remain unknown. Here, we present the first structural insights into hDicer in complex with tRNA, obtained by cryo-electron microscopy (cryo-EM), selective 2'-hydroxyl acylation analyzed by primer extension (SHAPE) and molecular dynamics (MD) simulations. Our results reveal that tRNAs adopt alternative conformations that are recognized and processed by hDicer. Furthermore, we show that tRNA cleavage by hDicer is facilitated by the m5C modification deposited by Nop2/SUN RNA methyltransferase 2 (NSUN2). Collectively, our findings redefine tRNAs as bona fide hDicer substrates and uncover a modification-dependent biogenetic pathway that reshapes the current understanding of the origins and regulation of human small RNAs. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=73 SRC="FIGDIR/small/744379v1_ufig1.gif" ALT="Figure 1000"> View larger version (24K): org.highwire.dtl.DTLVardef@ad78aborg.highwire.dtl.DTLVardef@cd3a7dorg.highwire.dtl.DTLVardef@1bb2594org.highwire.dtl.DTLVardef@1a0427e_HPS_FORMAT_FIGEXP M_FIG C_FIG

13
A generalisable method for the purification and biophysical characterisation of bacterial membrane receptors

Bosetto, F.; Zacharopoulou, M.; Bycroft, M.; Kish, M.; Zinzalla, G.; Rowling, P.; McLaughlin, S. H.; Phillips, J. J.; Itzhaki, L. S.; Mela, I.

2026-08-12 biophysics 10.64898/2026.08.12.744447 medRxiv
Top 0.4%
5.5%
Show abstract

Membrane-embedded bacterial receptors are challenging to express and purify in soluble form, yet their isolated domains are essential tools for structural and ligand-discovery studies. Pseudomonas aeruginosa relies on the TonB-dependent heme receptor HasR for iron acquisition, a process central to its pathogenicity. Here, we report a robust strategy for the recombinant expression, purification, and biophysical characterisation of the two soluble HasR domains directly involved in heme uptake: the N-terminal plug and the Secretin/TonB short N-terminal domain. Each domain was expressed individually in E. coli and purified to homogeneity, adopting well-folded conformations as confirmed by circular dichroism, NMR spectroscopy, and mass spectrometry. We then engineered a fusion construct containing both domains and systematically evaluated multiple solubilisation tags. A GST-His dual-affinity strategy enabled efficient purification of the construct, whereas His-tag alone resulted in insoluble protein and HLT-tag fusions suffered from non-specific proteolysis. Biophysical analyses revealed that the Secretin/TonB short N-terminal domain remains stably folded within the fusion construct, while the N-terminal plug domain becomes partially disordered, a finding further supported by hydrogen/deuterium exchange mass spectrometry. Together, these results establish a generalizable workflow for producing soluble receptor domains from membrane proteins and provide validated HasR constructs suitable for downstream ligand-screening applications, including aptamer and nanobody discovery.

14
Mutation of charged inner pore residues reduce E. coli β clamp residency and increase sliding rates on DNA

Liriano, M. L.; McCauley, M. J.; Ghosh, S.; Korzhnev, D.; Wales, T. E.; Williams, M. C.; Beuning, P. J.

2026-08-20 biochemistry 10.64898/2026.08.18.745641 medRxiv
Top 0.4%
5.5%
Show abstract

Sliding clamp proteins play central roles in DNA metabolism, including replication and repair. The ring-shaped E. coli beta clamp accommodates double-stranded DNA and serves as a platform for proteins involved in multiple DNA transactions. The inner pore of the beta clamp harbors a series of positively charged and polar residues that can bind to the negatively charged backbone of the DNA. These residues are arrayed so that they do not align with the charged phosphates of the DNA backbone. It is hypothesized that this arrangement of these residues provides for the movement of the clamp on DNA as it alternates which residues are bound to the DNA backbone. In this work, we mutated specific charged and polar residues that project into the inner pore of the beta clamp. The beta clamp variants are dimers and have similar thermal stability and in general a similar ability to complement a temperature sensitive strain for growth. One exception was beta-Q149A, which appeared as higher-order species on a native gel although its hydrogen-deuterium exchange pattern measured by mass spectrometry was overall similar to WT beta. These variants all had decreased binding to DNA after loading. Optical tweezers experiments were used to monitor loading on single DNA molecules and measure the rate of beta clamp sliding on DNA. Consistent with the hypothesized role of positively charged residues in the beta inner pore, mutation of one residue resulted in a faster rate of sliding on DNA.

15
PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets

Chi, L. A.; Ytreberg, F. M.

2026-08-18 biophysics 10.64898/2026.08.09.743757 medRxiv
Top 0.4%
5.5%
Show abstract

Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.

16
PyMOL plugin for Protein Circuit Topology

Dimins, M.; Bazba, A.; Mogyorosi, A.; Kennon, E.; Fiol, T. D.; Hagen, L. A.; Sheikhhassani, V.; Akulov, V.; Mashaghi, A.

2026-08-20 bioinformatics 10.1101/2025.10.21.683762 medRxiv
Top 0.4%
5.4%
Show abstract

Circuit Topology (CT) provides a fundamental framework for analysing folded polymer chains, with applications in functional annotation, protein engineering and drug development. We present a protein CT analysis plugin for PyMOL v3.1.6.1 with a graphical user interface (GUI), automatic installation, and novel features developed through integration with PyMOL's application programming interface (API). The plugin integrates various previously developed CT methodologies for studying structured proteins and their complexes as well as the dynamics of disordered proteins. Analysis of a representative protein and a molecular dynamics trajectory demonstrates the plugin's three analysis modes and their outputs. The plugin reproduces the reference ProteinCT implementation exactly on the structures tested, and is distributed with a versioned release, a pinned environment and a one-command reproduction of every result reported here.

17
Makeshift: a lightweight software for accessing and analyzing NMR data and protein dynamics

El Nesr, G.; Wayment-Steele, H. K.

2026-08-20 biophysics 10.64898/2026.08.17.745346 medRxiv
Top 0.4%
5.4%
Show abstract

Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.

18
NMR assignments and secondary structure analysis of the human 5MP1 C-terminal domain

Seker, A.; Anand, S.; Marintchev, A.

2026-08-18 biophysics 10.64898/2026.08.11.744028 medRxiv
Top 0.4%
5.4%
Show abstract

Eukaryotic translation initiation is tightly regulated by interactions among translation initiation factors (eIFs) that ensure accurate start codon selection. The translation regulator, eIF5 mimic protein 1 (5MP1) contributes to this process by competing with eIF5 for binding to eIF2, thereby increasing the stringency of translation initiation. Despite its important regulatory role and emerging involvement in tumorigenesis, structural information on human 5MP1 remains limited. Here, we report the near-complete backbone and partial side-chain NMR resonance assignments of the C-terminal domain of human 5MP1 (residues 250-419), carrying a W404E substitution that disrupts dimerization. The WT protein forms a dimer at NMR concentrations, which increases the effective size of the protein and also causes disappearance of peaks corresponding to aminoacids at the dimer interface due to conformational exchange. Backbone resonance assignments were completed for 96.4% of the non-proline residues. Secondary structure was analyzed using Chemical Shift Index (CSI) and compared with the AlphaFold structural model. Regions of disagreement between the experimental and computational secondary structure assignments were further examined using 15N-NOESY-HSQC spectra, allowing experimental validation of local structural features. While the AlphaFold model accurately reproduces the overall fold of the 5MP1 C-terminal domain, several localized discrepancies were identified, particularly near the N- and C-terminal regions of the domain, where experimental NMR data support alternative secondary structure assignments. These resonance assignments and experimentally validated structural features provide a foundation for future investigations of the molecular interactions, dynamics, and functions of 5MP1 in translation initiation.

19
The interaction between NC(p7)1-55 and p6 may regulate interactions with nucleic acids during assembly through modulation of Gag folding.

LARUE, V.; Nonin-Lecomte, S.

2026-09-01 biophysics 10.64898/2026.08.28.747767 medRxiv
Top 0.4%
5.3%
Show abstract

We present the solution structures of HIV-1 proteins NC(p7)1-55 corresponding to the full-length NC(p7) and mature p6. The studies were carried in water and, to mimic the membrane, in micellar DPC (Dodecylphosphocholine) conditions. Our results unravel for the first time the structure adopted by the N-terminal amino acids of the free NC(p7)1-55, with the formation of a small helix spanning residues F6 to R10. Our NMR and Fluorescence Anisotropy data disclose an interaction between NC(p7)1-55 and p6 both in water and DPC, with respective Kd of 2.5mM and 370 mM at 23{degrees}C. The interaction is thus strengthened in lipidic conditions. Protein p6 stabilizes the N-terminus of NC(p7)1-55 while increasing at the same time the dynamic of the first zinc finger. Although the entire p6 sequence is involved in the interaction, we show that its C-terminal region is particularly sensitive to the presence of NC(p7)1-55, with a propensity of forming a a helix ranging from amino acids S111 to F116. This study brings experimental evidence of a direct protein-protein interaction between p6 and the N-terminal region of NC(p7)1-55. We further show that such interaction is readily accommodated within the NC(p15) framework and hypothesize that it may facilitate the selective assembly of assembly of the viral genomic RNA (gRNA) in the cell.

20
SilkRoute: A Descriptor-Driven Framework for Reproducible Multi-Source Biomolecular Data Acquisition

Fernandez, D.; Garcia-Vinuesa, J.; Alvarez-Saravia, D.; Soto-Garcia, M.; Medina-Franco, J. L.; Sepulveda-Yanez, J.; Cadet, X.; Cadet, F.; Davari, M. D.; Uribe-Paredes, R.; Herrera-Rocha, F.; Medina-Ortiz, D.

2026-08-18 bioinformatics 10.64898/2026.08.11.744100 medRxiv
Top 0.4%
5.2%
Show abstract

BackgroundBiomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedures difficult to inspect, reproduce, or adapt across studies. We present SilkRoute, an open-source Python framework that formalizes biomolecular data acquisition as descriptor-defined, source-aware, and provenance-tracked workflows, providing a reproducible foundation for multi-source biomolecular dataset construction. ResultsSilkRoute uses machine-readable YAML descriptors to specify dataset intent, biomolecular modality, workflow mode, query logic, enrichment resources, execution parameters, and export settings. These descriptors drive a common execution model that coordinates primary retrieval and downstream enrichment while preserving source-specific outputs, interaction evidence when available, the original workflow configuration, metadata, and run summaries. We evaluated this model through three representative acquisition scenarios spanning proteins, compounds, and molecular interactions. In the protein-centered workflow, SilkRoute retrieved 2,444 reviewed antimicrobial protein records from UniProt and generated complementary outputs from AlphaFold DB, InterPro, Pathway Commons, and the Protein Data Bank. In the compound-centered workflow, a ChEMBL IC50 query produced 1,445,939 activity records organized into query-defined potency ranges. In the interaction-centered workflow, 2,253 UniProt protein records were expanded with 902,713 BioGRID interaction records and 5,702 STRING interaction-partner records. Across these scenarios, the framework successfully applied the same descriptor-defined acquisition model to distinct biomolecular entity types, retrieval strategies, enrichment paths, and output structures. ConclusionsSilkRoute extends beyond sequence retrieval by providing a reusable acquisition layer for constructing multi-source biomolecular datasets. By separating primary retrieval from enrichment and preserving source-aware outputs together with workflow descriptors and execution metadata, the framework makes acquisition procedures easier to inspect, reproduce, archive, and adapt. SilkRoute does not replace biological curation, label validation, deduplication, partitioning, or benchmarking, but provides structured and traceable acquisition packages that support these downstream processes.